Skip to content

[Klaud Cold] Update kimik3-fp4-gb200-dynamo-vllm-agentic-dspark-mooncake-dcp16-agg vLLM image to nightly-dev-arm64-cu13-3696c77 (digest-pinned) / 将 kimik3-fp4-gb200-dynamo-vllm-agentic-dspark-mooncake-dcp16-agg 的 vLLM 镜像更新至 nightly-dev-arm64-cu13-3696c77(按 digest 固定) - #2956

Closed
Klaud-Cold wants to merge 4 commits into
mainfrom
klaud/auto-51d9acec3b181791-8eede55d6ccb7552

Conversation

@Klaud-Cold

Copy link
Copy Markdown
Collaborator

Refresh the kimik3-fp4-gb200-dynamo-vllm-agentic-dspark-mooncake-dcp16-agg AgentX family from the locally built dev image vllm/vllm-openai:nightly-dev-arm64-cu13.0.1-75c2eef (vLLM dev commit 75c2eef, 2026-08-14) to the newer same-lineage dev image vllm/vllm-openai:nightly-dev-arm64-cu13-3696c77 (vLLM dev commit 3696c77, 2026-09-01), pinned by digest because the tag is mutable. Model, TP16/DCP16 topology, DSpark K=4 speculation, Mooncake DRAM offload, workloads, Dynamo pin and all five concurrency points are unchanged.

Baseline

  • Published date: 2026-08-21 (benchmarks?model=Kimi-K3&date=2026-08-21&exact=true, workflow-info?date=2026-08-21&benchmarkType=agentic_traces)
  • Old image: vllm/vllm-openai:nightly-dev-arm64-cu13.0.1-75c2eef (registry manifest today: sha256:be4945a0…73f60d, created 2026-08-17; the tag was re-pushed on 2026-09-03, so the exact bytes benchmarked on 2026-08-21 cannot be re-verified)
  • Workload/topology: AgentX agentic-coding traces, Kimi-K3 FP4, aggregated TP16/DCP16 on 4× GB200 nodes (16 GPUs), DSpark K=4 with synthetic acceptance length 3.36, Mooncake DRAM KV offload (96 GB), max-num-seqs 2
  • Producer: run 32424103771 attempt 2 at head 349aa162879bdae1c39ca71282f5f555c8aa2946 (PR #2700); recipe fingerprint a295bb2c32…
  • Source links: old vLLM dev branch vs main 925ea7e…75c2eef (18 dev-only DSpark/DCP/Mooncake fixes); new vLLM dev branch vs main 080a66a…3696c77 (32 dev commits); Dynamo pin ba83080 unchanged; srt-slurm v1.0.50 unchanged
conc tput/GPU (tok/s) output tok/s median TTFT (s) median ITL (s) median interactivity (tok/s)
1 917.30 112.26 2.158 0.00448 223.2
2 973.96 131.38 2.078 0.00456 219.3
4 1489.05 164.49 2.070 0.00479 208.8
8 2161.78 220.23 15.369 0.00509 196.5
16 2300.06 236.18 48.465 0.00515 194.2

Published eval (2026-08-21, TP16/DCP16 MTP, conc 16): gsm8k em_strict 0.9727 (n=1319, run 32424103771). The current generator selects the kimi-vendor / kimi_tool_call_schema eval for this family instead of gsm8k, so eval values are N/A for direct comparison.


kimik3-fp4-gb200-dynamo-vllm-agentic-dspark-mooncake-dcp16-agg AgentX 配置的镜像从本地构建的开发镜像 vllm/vllm-openai:nightly-dev-arm64-cu13.0.1-75c2eef(vLLM 开发分支提交 75c2eef,2026-08-14)更新为同一开发谱系的较新镜像 vllm/vllm-openai:nightly-dev-arm64-cu13-3696c77(vLLM 开发分支提交 3696c77,2026-09-01),并因标签可变而按 digest 固定。模型、TP16/DCP16 拓扑、DSpark K=4 投机解码、Mooncake DRAM 卸载、工作负载、Dynamo 版本和全部五个并发点均保持不变。

基线

  • 发布日期:2026-08-21(benchmarks?model=Kimi-K3&date=2026-08-21&exact=trueworkflow-info?date=2026-08-21&benchmarkType=agentic_traces
  • 旧镜像:vllm/vllm-openai:nightly-dev-arm64-cu13.0.1-75c2eef(当前仓库 manifest:sha256:be4945a0…73f60d,构建于 2026-08-17;该标签在 2026-09-03 被重新推送,因此无法再核实 2026-08-21 实际测试的镜像字节)
  • 工作负载/拓扑:AgentX agentic-coding 轨迹,Kimi-K3 FP4,4 个 GB200 节点(16 GPU)上的聚合 TP16/DCP16,DSpark K=4(合成接受长度 3.36),Mooncake DRAM KV 卸载(96 GB),max-num-seqs 2
  • 产出运行:run 32424103771 attempt 2,head 349aa162879bdae1c39ca71282f5f555c8aa2946PR #2700);配方指纹 a295bb2c32…
  • 源码链接:旧 vLLM 开发分支与 main 对比 925ea7e…75c2eef(18 个仅在开发分支的 DSpark/DCP/Mooncake 修复);新 vLLM 开发分支与 main 对比 080a66a…3696c77(32 个开发提交);Dynamo 固定版本 ba83080 不变;srt-slurm v1.0.50 不变
并发 每 GPU 吞吐(tok/s) 输出 tok/s TTFT 中位数(s) ITL 中位数(s) 交互性中位数(tok/s)
1 917.30 112.26 2.158 0.00448 223.2
2 973.96 131.38 2.078 0.00456 219.3
4 1489.05 164.49 2.070 0.00479 208.8
8 2161.78 220.23 15.369 0.00509 196.5
16 2300.06 236.18 48.465 0.00515 194.2

已发布评测(2026-08-21,TP16/DCP16 MTP,并发 16):gsm8k em_strict 0.9727(n=1319,run 32424103771)。当前生成器为该配置选择 kimi-vendor / kimi_tool_call_schema 评测而非 gsm8k,因此评测值无法直接对比,记为 N/A。

🤖 Generated with Claude Code

… to nightly-dev-arm64-cu13-3696c77

Move kimik3-fp4-gb200-dynamo-vllm-agentic-dspark-mooncake-dcp16-agg and its
unshared srt-slurm recipe from vllm/vllm-openai:nightly-dev-arm64-cu13.0.1-75c2eef
to vllm/vllm-openai:nightly-dev-arm64-cu13-3696c77, pinned by digest
sha256:42e17a3c600c043c45e0623e4568cd4d5903dd76525a85f5df6ca6fa65358b12.
model.container and identity.container.image match the master image; topology,
speculation, workloads, Dynamo pin and concurrency points are unchanged.

将 kimik3-fp4-gb200-dynamo-vllm-agentic-dspark-mooncake-dcp16-agg 及其未共享的
srt-slurm 配方镜像从 vllm/vllm-openai:nightly-dev-arm64-cu13.0.1-75c2eef 更新为
vllm/vllm-openai:nightly-dev-arm64-cu13-3696c77,并按 digest
sha256:42e17a3c600c043c45e0623e4568cd4d5903dd76525a85f5df6ca6fa65358b12 固定。
model.container 与 identity.container.image 与主配置镜像一致;拓扑、投机解码、
工作负载、Dynamo 版本和并发点均保持不变。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@Klaud-Cold

Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

Initial attempt

  • Image: vllm/vllm-openai:nightly-dev-arm64-cu13-3696c77@sha256:42e17a3c600c043c45e0623e4568cd4d5903dd76525a85f5df6ca6fa65358b12 (arm64, CUDA 13.0.1, TORCH_CUDA_ARCH_LIST includes 10.0+PTX; built 2026-09-02 from vLLM dev commit 3696c77)
  • PR head: 58c3e91617f47c7817903024616aae41e4521183
  • Change: master image plus recipe model.container / identity.container.image only. Generated matrix at this head differs from main only in image: 5 points (conc 1/2/4/8/16), TP16/DCP16, 4 nodes, DSpark K=4, Mooncake DRAM offload, kimi_tool_call_schema eval per point.
  • Targeted smoke: test-config --config-files configs/nvidia-master.yaml --config-keys kimik3-fp4-gb200-dynamo-vllm-agentic-dspark-mooncake-dcp16-agg --trim-conc → conc 1 + its eval. Run: 34446587513 (dispatched 2026-09-10 06:44 UTC from main, fail-fast=true, klaud-run=true).

Upstream source comparison

  • Old image commit 75c2eef (2026-08-14) sits on a vLLM dev branch 18 commits ahead of main (compare); those commits are DSpark-under-DCP, Mooncake-store-DCP-on-hybrid-models, TokenspeedMLA and KV-cache fixes that this DCP16 + Mooncake + DSpark recipe exercises. The registry image carries VLLM_BUILD_PIPELINE=local and no build commit label, so provenance is inferred from the tag suffix only.
  • New image commit 3696c77 (2026-09-01) sits on the successor dev branch: main at 080a66a (2026-08-26) plus 32 commits (compare) including "Support MooncakeStore with hybrid DCP prefix caching", "fix k3 hybrid mooncake recompute handling", "Enable TokenSpeed MLA for interleaved DCP speculation", "Honor DSpark draft load config", "Fix DFlash profiling inputs for DCP", Mooncake load/put hardening and Mamba/KDA replay fixes. This exact image is already the shipped image for the GB300 Kimi-K3 Mooncake DCP8 families (published 2026-09-04) with the same Dynamo pin ba83080 and srt-slurm v1.0.50, so no launch-path or runtime patch is needed.
  • Flag/config check at 3696c77: every option the recipe uses (TOKENSPEED_MLA, dspark speculative method with draft_sample_method/rejection_sample_method, decode-context-parallel-size, dcp-comm-backend, prefix-match-unit, safetensors-load-strategy, moe-backend, enable-cumem-allocator, kv-cache-memory, MooncakeStoreConnector, VLLM_USE_DIRECT_DCP_*, VLLM_MOONCAKE_LOAD_RECV_THREADS, VLLM_PREFIX_CACHE_RETENTION_INTERVAL, VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS) is present in the parsers with the same names; no renamed or removed option found. VLLM_ENABLE_K3_LATENT_MOE_TAIL_FUSION is not declared in vllm/envs.py at either commit, so it is left as-is.
  • Release hint v0.29.0 (98dff2a, image published 2026-09-09) was evaluated and not selected: the release branch predates upstream #55472 (DSpark DCP fix, merged 2026-09-08) and contains only 106 of the 568 non-test lines added by the old dev branch (none of the Mooncake-store DCP, KV-cache-manager or Kimi-K3 KDA changes), so it is not a drop-in replacement for this recipe.
  • The old tag nightly-dev-arm64-cu13.0.1-75c2eef was re-pushed on 2026-09-03 (manifest created 2026-08-17); the new tag is therefore pinned by digest.

Results

  • Run 34446587513: failed (benchmark job and eval-only job both failed after ~25 min; no benchmark point, no eval score → all deltas N/A).
  • First server error (frontend, watchtower-navy-cn01_frontend_0.out): after installing Dynamo ba83080 from source, python3 -m dynamo.frontend --dyn-chat-processor vllm … exits with ModuleNotFoundError: No module named 'vllm.entrypoints.openai.cli_args' (Flag '--chat-processor vllm' requires vllm be installed). The 16 vLLM workers had reached torch.distributed init (world_size=16) with no engine error before the frontend crash tore the job down.
  • Diagnosis: vLLM moved cli_args.py out of vllm/entrypoints/openai/ in vllm-project/vllm#53659 (2026-08-25), so it is absent from every image built after that date (3696c77, v0.29.0). Dynamo's vLLM chat-processor path imports the old location at the pinned ba83080 and still on Dynamo main, so no Dynamo pin fixes it. The GB300 Kimi-K3 recipes on the same image and Dynamo pin do not set dyn-chat-processor: vllm, which is why they run.
  • Next step: Repair 1/5 switches this recipe's frontend to the native Dynamo chat processor (see the Repair 1/5 comment).

初始尝试

  • 镜像:vllm/vllm-openai:nightly-dev-arm64-cu13-3696c77@sha256:42e17a3c600c043c45e0623e4568cd4d5903dd76525a85f5df6ca6fa65358b12(arm64,CUDA 13.0.1,TORCH_CUDA_ARCH_LIST10.0+PTX;2026-09-02 基于 vLLM 开发分支提交 3696c77 构建)
  • PR head:58c3e91617f47c7817903024616aae41e4521183
  • 变更:仅主配置 image 以及配方的 model.container / identity.container.image。该 head 生成的矩阵与 mainimage 不同:5 个点(并发 1/2/4/8/16),TP16/DCP16,4 节点,DSpark K=4,Mooncake DRAM 卸载,每点附带 kimi_tool_call_schema 评测。
  • 定向冒烟:test-config --config-files configs/nvidia-master.yaml --config-keys kimik3-fp4-gb200-dynamo-vllm-agentic-dspark-mooncake-dcp16-agg --trim-conc → 并发 1 及其评测。运行:34446587513(2026-09-10 06:44 UTC 从 main 分发,fail-fast=trueklaud-run=true)。

上游源码对比

  • 旧镜像提交 75c2eef(2026-08-14)位于领先 main 18 个提交的 vLLM 开发分支(对比);这些提交是本 DCP16 + Mooncake + DSpark 配方所依赖的 DSpark-DCP、Mooncake 存储对混合模型 DCP、TokenspeedMLA 与 KV 缓存修复。仓库镜像带有 VLLM_BUILD_PIPELINE=local 且没有构建提交标签,来源仅能从标签后缀推断。
  • 新镜像提交 3696c77(2026-09-01)位于后继开发分支:main080a66a(2026-08-26)加 32 个提交(对比),包括 “Support MooncakeStore with hybrid DCP prefix caching”、“fix k3 hybrid mooncake recompute handling”、“Enable TokenSpeed MLA for interleaved DCP speculation”、“Honor DSpark draft load config”、“Fix DFlash profiling inputs for DCP”、Mooncake 加载/写入加固以及 Mamba/KDA 回放修复。该镜像已是 GB300 Kimi-K3 Mooncake DCP8 配置的正式镜像(2026-09-04 发布),使用相同的 Dynamo 版本 ba83080 与 srt-slurm v1.0.50,无需任何启动路径或运行时补丁。
  • 3696c77 上的参数/配置检查:配方使用的全部选项(TOKENSPEED_MLA、带 draft_sample_method/rejection_sample_methoddspark 投机方法、decode-context-parallel-sizedcp-comm-backendprefix-match-unitsafetensors-load-strategymoe-backendenable-cumem-allocatorkv-cache-memoryMooncakeStoreConnectorVLLM_USE_DIRECT_DCP_*VLLM_MOONCAKE_LOAD_RECV_THREADSVLLM_PREFIX_CACHE_RETENTION_INTERVALVLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS)在解析器中同名存在,未发现重命名或移除的选项。VLLM_ENABLE_K3_LATENT_MOE_TAIL_FUSION 在两个提交的 vllm/envs.py 中均未声明,保持原样。
  • 版本提示 v0.29.098dff2a,镜像 2026-09-09 发布)经评估未采用:其发布分支早于上游 #55472(DSpark DCP 修复,2026-09-08 合并),且仅包含旧开发分支新增的 568 行非测试代码中的 106 行(不含 Mooncake 存储 DCP、KV 缓存管理器和 Kimi-K3 KDA 变更),因此不能直接替换本配方。
  • 旧标签 nightly-dev-arm64-cu13.0.1-75c2eef 于 2026-09-03 被重新推送(manifest 创建于 2026-08-17),因此新标签按 digest 固定。

结果

  • 运行 34446587513失败(基准任务与仅评测任务均在约 25 分钟后失败;无基准点、无评测分数,所有差值记为 N/A)。
  • 首个服务端错误(前端,watchtower-navy-cn01_frontend_0.out):从源码安装 Dynamo ba83080 后,python3 -m dynamo.frontend --dyn-chat-processor vllm …ModuleNotFoundError: No module named 'vllm.entrypoints.openai.cli_args'Flag '--chat-processor vllm' requires vllm be installed)退出。前端崩溃导致作业被拆除之前,16 个 vLLM worker 已完成 torch.distributed 初始化(world_size=16),未出现引擎错误。
  • 诊断:vLLM 在 vllm-project/vllm#53659(2026-08-25)中将 cli_args.py 移出 vllm/entrypoints/openai/,因此该日期之后构建的所有镜像(3696c77v0.29.0)都不再包含它。Dynamo 的 vLLM 聊天处理器路径在固定版本 ba83080 以及 Dynamo main 上仍导入旧路径,因此更换 Dynamo 版本无法解决。使用相同镜像和 Dynamo 版本的 GB300 Kimi-K3 配方未设置 dyn-chat-processor: vllm,因此能够运行。
  • 下一步:修复 1/5 将本配方的前端切换到 Dynamo 原生聊天处理器(见修复 1/5 评论)。

…P16 DSpark

vLLM moved vllm/entrypoints/openai/cli_args.py (vllm-project/vllm#53659,
2026-08-25), which Dynamo's `--dyn-chat-processor vllm` frontend path imports
at the pinned ba83080 and on Dynamo main. Drop dyn-chat-processor and the
vLLM-parser-only frontend flags (tool-call-parser, reasoning-parser,
enable-auto-tool-choice) so the frontend uses Dynamo's native processor, as
the GB300 Kimi-K3 recipes already do. Tool-call and reasoning parsing remain
kimi_k3 through the workers' dyn-tool-call-parser / dyn-reasoning-parser.

vLLM 在 vllm-project/vllm#53659(2026-08-25)中移动了
vllm/entrypoints/openai/cli_args.py,而 Dynamo 固定版本 ba83080 及 main 上的
`--dyn-chat-processor vllm` 前端路径仍导入该模块。移除 dyn-chat-processor 及仅
vLLM 解析器接受的前端参数(tool-call-parser、reasoning-parser、
enable-auto-tool-choice),使前端使用 Dynamo 原生处理器,与 GB300 Kimi-K3 配方
一致。工具调用与推理解析仍通过 worker 的 dyn-tool-call-parser /
dyn-reasoning-parser 使用 kimi_k3。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Klaud-Cold

Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

Repair 1/5

  • Image: unchanged, vllm/vllm-openai:nightly-dev-arm64-cu13-3696c77@sha256:42e17a3c600c043c45e0623e4568cd4d5903dd76525a85f5df6ca6fa65358b12
  • PR head: a090cd67522d8062408c79f8e90a276f2aeb615a
  • Change (unshared recipe agg-gb200-dcp16-dspark4-maxseq2-mooncake-agentic.yaml, frontend.args only): drop dyn-chat-processor: "vllm" and the vLLM-parser-only flags tool-call-parser, reasoning-parser, enable-auto-tool-choice; keep trust-remote-code, router-mode: random, router-session-affinity-ttl-secs: 900. The frontend now uses Dynamo's native chat processor, the same configuration the published GB300 Kimi-K3 recipes use on this image and Dynamo pin. Tool-call and reasoning parsing remain kimi_k3 via the workers' existing dyn-tool-call-parser / dyn-reasoning-parser registration (Dynamo ba83080 components/src/dynamo/vllm/main.py lines 710-711). Model, topology, DSpark, Mooncake, workloads, benchmark command, resources and Dynamo pin are unchanged.
  • Outcome: targeted smoke (conc 1 benchmark + eval) passed; next step is the final full sweep (see the Final full sweep comment).
  • Why not another Dynamo pin: Dynamo main (components/src/dynamo/frontend/main.py) still imports vllm.entrypoints.openai.cli_args, so no pin restores the vLLM chat-processor path on a post-2026-08-25 vLLM; patching Dynamo or vLLM is out of scope.
  • Note for reviewers: request pre/post-processing moves from vLLM's Python OpenAI processor to Dynamo's Rust frontend. This is a serving-path change, not an engine change; GB300 Kimi-K3 gsm8k evals on this path (2026-09-04, 0.965–0.973) match the GB200 vLLM-processor baseline (0.9727).
  • Targeted smoke: same test-config … --trim-conc command. Run: 34449828534 (dispatched 2026-09-10 07:25 UTC from main, fail-fast=true, klaud-run=true).

Results

  • Run 34449828534: benchmark job passed (conc 1, TP16/DCP16, 4 nodes, 3594 s). The eval-only job failed once at Dynamo install on watchtower-navy-cn04 (pip … ResolutionImpossible … pydantic-core: no matching distributions available for your environment, a package-fetch failure on that node; the other three nodes installed the same cached wheel) and was re-run as attempt 2 of the same run (no code change) and passed its score gate: kimi_tool_call_schema smoke 0.5 (1 of 2 cases; the non-stream case passed, the stream case failed with streamed response had no tool_calls (chunks=17, finish_reason=stop, content=)). No published kimi_tool_call_schema result exists for this or any family, so the eval delta is N/A; the published 2026-08-21 eval for this family was gsm8k (0.9727), which the current generator no longer selects. The streaming-mode tool-call loss is consistent with the native kimi_k3 parser warnings seen in the benchmark job and is flagged as a serving-path regression for reviewers.
  • Smoke point vs published baseline (conc 1, same recipe; smoke only, not a curve):
conc 1 baseline (75c2eef, 2026-08-21) smoke (3696c77, run 34449828534) delta
total tput / GPU (tok/s) 917.30 955.51 +4.2%
output tok/s 112.26 117.33 +4.5%
input tok/s 14564.6 15170.8 +4.2%
median TTFT (s) 2.158 1.587 -26.5%
median ITL (s) 0.00448 0.00433 -3.3%
median interactivity (tok/s) 223.2 230.7 +3.4%
profiled requests 232 234 +2
profiling-phase invalid results 1 / 233 (0.43%) 7 / 241 (2.9%) +6
  • The invalid results are AIPerf InvalidInferenceResultError ("no responses with actual content"): both runs drop 10 of 11 warmup requests this way; in profiling the new run has 7 (six of them in one conversation during the first 18 minutes, all answered status=success in ~1–2 s by the frontend). One of the seven coincides with the native kimi_k3 tool parser dropping an incomplete call (dynamo_parsers … kimi_k3_incomplete_call, 74 such warnings across 11 requests in the run). Reported for reviewer awareness; the point stays within the 25% failed-request threshold and the excluded requests do not enter the throughput/latency metrics above.
  • Server-side: GPU prefix-cache hit rate 0.961 (baseline 0.871), Mooncake CPU hit rate 0.005 (baseline 0.030); no engine errors in the worker logs (the TCPStore sendBytes traces appear only at teardown, 09:01 UTC).

修复 1/5

  • 镜像:不变,vllm/vllm-openai:nightly-dev-arm64-cu13-3696c77@sha256:42e17a3c600c043c45e0623e4568cd4d5903dd76525a85f5df6ca6fa65358b12
  • PR head:a090cd67522d8062408c79f8e90a276f2aeb615a
  • 变更(仅未共享配方 agg-gb200-dcp16-dspark4-maxseq2-mooncake-agentic.yamlfrontend.args):移除 dyn-chat-processor: "vllm" 以及仅 vLLM 解析器接受的 tool-call-parserreasoning-parserenable-auto-tool-choice;保留 trust-remote-coderouter-mode: randomrouter-session-affinity-ttl-secs: 900。前端改用 Dynamo 原生聊天处理器,与已发布的 GB300 Kimi-K3 配方在同一镜像和 Dynamo 版本上的配置一致。工具调用与推理解析仍通过 worker 现有的 dyn-tool-call-parser / dyn-reasoning-parser 注册使用 kimi_k3(Dynamo ba83080 components/src/dynamo/vllm/main.py 第 710-711 行)。模型、拓扑、DSpark、Mooncake、工作负载、基准命令、资源和 Dynamo 版本均不变。
  • 结论:定向冒烟(并发 1 基准 + 评测)通过;下一步为最终完整扫描(见“最终完整扫描”评论)。
  • 为何不更换 Dynamo 版本:Dynamo maincomponents/src/dynamo/frontend/main.py)仍导入 vllm.entrypoints.openai.cli_args,因此在 2026-08-25 之后的 vLLM 上没有任何版本能恢复 vLLM 聊天处理器路径;修补 Dynamo 或 vLLM 超出范围。
  • 审阅提示:请求的前后处理从 vLLM 的 Python OpenAI 处理器转移到 Dynamo 的 Rust 前端。这是服务路径变更而非引擎变更;GB300 Kimi-K3 在该路径上的 gsm8k 评测(2026-09-04,0.965–0.973)与 GB200 vLLM 处理器基线(0.9727)相当。
  • 定向冒烟:相同的 test-config … --trim-conc 命令。运行:34449828534(2026-09-10 07:25 UTC 从 main 分发,fail-fast=trueklaud-run=true)。

结果

  • 运行 34449828534:基准任务通过(并发 1,TP16/DCP16,4 节点,3594 秒)。仅评测任务在 watchtower-navy-cn04 上因 Dynamo 安装失败一次(pip … ResolutionImpossible … pydantic-core: no matching distributions available for your environment,该节点的软件包获取失败;另外三个节点安装了同一缓存 wheel),已在同一运行中作为第 2 次尝试重跑(无代码变更)并通过分数门槛:kimi_tool_call_schema 冒烟得分 0.5(2 个用例中 1 个通过;非流式用例通过,流式用例失败,报 streamed response had no tool_calls (chunks=17, finish_reason=stop, content=))。该配置及任何配置均无已发布的 kimi_tool_call_schema 结果,评测差值记为 N/A;该配置 2026-08-21 发布的评测为 gsm8k(0.9727),当前生成器已不再选择该评测。流式模式下工具调用丢失与基准任务中原生 kimi_k3 解析器的警告一致,已作为服务路径回归标记供审阅者参考。
  • 冒烟点与已发布基线对比(并发 1,配方相同;仅为冒烟,非完整曲线):
并发 1 基线(75c2eef,2026-08-21) 冒烟(3696c77,运行 34449828534) 差值
每 GPU 总吞吐(tok/s) 917.30 955.51 +4.2%
输出 tok/s 112.26 117.33 +4.5%
输入 tok/s 14564.6 15170.8 +4.2%
TTFT 中位数(s) 2.158 1.587 -26.5%
ITL 中位数(s) 0.00448 0.00433 -3.3%
交互性中位数(tok/s) 223.2 230.7 +3.4%
采样请求数 232 234 +2
采样阶段无效结果 1 / 233(0.43%) 7 / 241(2.9%) +6
  • 无效结果均为 AIPerf 的 InvalidInferenceResultError(“未收到包含实际内容的响应”):两次运行都以此丢弃 11 个预热请求中的 10 个;采样阶段新运行有 7 个(其中 6 个集中在前 18 分钟的同一会话,前端均在约 1–2 秒内以 status=success 返回)。其中 1 个与原生 kimi_k3 工具解析器丢弃不完整调用(dynamo_parsers … kimi_k3_incomplete_call,整个运行共 74 条此类警告,涉及 11 个请求)同时发生。此处如实记录供审阅者参考;该点仍在 25% 失败请求阈值之内,被排除的请求不计入上表吞吐/延迟指标。
  • 服务端:GPU 前缀缓存命中率 0.961(基线 0.871),Mooncake CPU 命中率 0.005(基线 0.030);worker 日志无引擎错误(TCPStore sendBytes 堆栈仅出现在 09:01 UTC 拆除阶段)。

…refresh

Append the perf-changelog entry for PR #2956: image bump to
vllm/vllm-openai:nightly-dev-arm64-cu13-3696c77 (digest-pinned) and the
native Dynamo chat processor for the frontend.

为 PR #2956 追加 perf-changelog 条目:镜像升级为
vllm/vllm-openai:nightly-dev-arm64-cu13-3696c77(按 digest 固定),前端改用
Dynamo 原生聊天处理器。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Take main's perf-changelog.yaml verbatim and re-append this PR's entry at the
tail so the PR is mergeable and run-sweep can trigger.

合并 main:按原样采用 main 的 perf-changelog.yaml,并在末尾重新追加本 PR 的
条目,使 PR 可合并并能触发 run-sweep。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Klaud-Cold

Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

Final full sweep

  • PR head: 256689714a9b4e405dc627c9419cd92499ea8dc1 (merge of main at 9199e34f plus the perf-changelog entry; image and recipe unchanged from Repair 1/5). The first changelog head e9e068ac was CONFLICTING against main (two changelog entries landed on main after this branch's base), which suppresses pull_request-triggered workflows, so no run-sweep run was created; full-sweep-enabled was removed, main merged with main's perf-changelog.yaml taken verbatim and this PR's entry re-appended at the tail, then the label re-applied. No repair budget used for this sync.
  • Trigger: full-sweep-enabled re-applied at 09:36 UTC. run-sweep runs on this head: 34461597759 (synchronize) and 34461599480 (labeled); the concurrency group cancelled the synchronize run before any benchmark started; 34461599480 is the final sweep run.
  • Scope: the complete family matrix from the changelog entry: 5 benchmark points (conc 1/2/4/8/16, TP16/DCP16, 4 nodes each) plus the default kimi_tool_call_schema eval per point. No trimming or eval-selection modifiers.

Results

  • 10:20 UTC: benchmark points c1/c2/c4/c8 running in parallel since 09:40 UTC; c16 and the five eval-only jobs queued behind them. No failures so far.
  • 11:00 UTC: c1/c2/c4/c8 still running (80 min in, inside the 60-minute profiling window after model load); c16 and the evals still queued.
  • 11:35 UTC: c1/c2/c4/c8 benchmark points passed; c16 and the evals for c2/c4/c16 are running, evals for c1/c8 queued. Interim per-point comparison against the published 2026-08-21 baseline (same recipe; final values follow ingestion):
conc total tput / GPU (tok/s) output tok/s median TTFT (s) median ITL (s) median interactivity (tok/s) profiled requests
1 917.3 → 971.0 (+5.8%) 112.3 → 117.4 (+4.6%) 2.158 → 1.589 (-26.4%) 0.00448 → 0.00428 (-4.5%) 223.2 → 233.6 (+4.7%) 232 → 238
2 974.0 → 1036.2 (+6.4%) 131.4 → 142.9 (+8.8%) 2.078 → 1.345 (-35.3%) 0.00456 → 0.00434 (-4.8%) 219.3 → 230.5 (+5.1%) 372 → 390
4 1489.0 → 1553.7 (+4.3%) 164.5 → 170.4 (+3.6%) 2.070 → 1.408 (-32.0%) 0.00479 → 0.00444 (-7.3%) 208.8 → 225.3 (+7.9%) 507 → 517
8 2161.8 → 2320.8 (+7.4%) 220.2 → 244.7 (+11.1%) 15.369 → 7.730 (-49.7%) 0.00509 → 0.00478 (-6.1%) 196.5 → 209.1 (+6.4%) 840 → 900

Invalid-result accounting in these four points: c1 11/252 (all warmup), c2 24/419 (22 warmup), c4 44/568 (all warmup), c8 68/1000 (warmup 90) — the profiling-phase "no content" responses seen in the smoke did not recur beyond c2's two.

  • 11:55 UTC: final sweep failed at the eval gate. The run's default eval for this family is lm-eval gsm8k per point (not the kimi_tool_call_schema smoke that the trimmed targeted run used). Three evals finished and all failed Verify eval scores (threshold 0.90 strict-match):
conc gsm8k strict (published baseline 0.9727 at c16) gsm8k flexible (baseline 0.9719) responses without a #### N answer line
2 0.8165 0.9090 N/A (samples not inspected)
4 0.7991 0.9030 N/A (samples not inspected)
16 0.8120 0.9067 219 / 1319 (baseline run: 0 / 1319)
  • Diagnosis (first server-side cause): the eval server was healthy and all 1319 requests completed, but 219 responses at c16 end mid-sentence before the answer line (median response length 258 vs 293 characters in the baseline samples; no empty responses; 29 genuinely wrong answers vs 36 in the baseline). This is the Dynamo native chat processor at pin ba83080 (Repair 1/5) truncating response tails: its kimi_k3 XTML tool parser buffers trailing text and drops it as an "incomplete call" at stream end (dynamo_parsers::tool_calling::xtml::kimi_k3_parser … kimi_k3_incomplete_call, also seen in the benchmark frontend logs), and at ba83080 the parser runs even for requests without tools. The vLLM chat-processor path used by the baseline cannot run on this image (the vllm.entrypoints.openai.cli_args import removed upstream on 2026-08-25, see Initial attempt), and Dynamo main still imports the same path.
  • Benchmark side: c1/c2/c4/c8 passed with the per-point deltas above; c16 benchmark and the c1/c8 evals were still running at the stop and are cancelled by finish because the sweep can no longer validate on this head.
  • Stop decision: no runtime patch is permitted, so the remaining in-scope option would be moving the recipe's Dynamo pin to a commit that scopes tool parsing to requests that declare tools (fix(frontend): scope tool parsing to request semantics ai-dynamo/dynamo#13359, merged 2026-08-20, or later) and re-validating worker compatibility with vLLM 3696c77. That needs a new targeted smoke (~1.7 h) plus a full sweep (~3 h), which does not fit the remaining session budget, so this candidate stops as failed after 1 of 5 repairs with the branch retained for manual review. Recommended follow-up for maintainers: bump dynamo.hash / identity.frameworks.dynamo in agg-gb200-dcp16-dspark4-maxseq2-mooncake-agentic.yaml alongside this image, or wait for an upstream Dynamo frontend fix for the new vLLM cli_args location so dyn-chat-processor: vllm can be restored.

最终完整扫描

  • PR head:256689714a9b4e405dc627c9419cd92499ea8dc1(合并 main9199e34f 并包含 perf-changelog 条目;镜像与配方与修复 1/5 一致)。首个变更日志 head e9e068acmain 存在冲突(CONFLICTING,此分支基线之后 main 上新增了两条变更日志条目),GitHub 因此不触发 pull_request 工作流,未生成 run-sweep 运行;已移除 full-sweep-enabled,合并 main(按原样采用 mainperf-changelog.yaml 并在末尾重新追加本 PR 条目),然后重新添加标签。此次同步不消耗修复预算。
  • 触发:09:36 UTC 重新添加 full-sweep-enabled。该 head 上的 run-sweep 运行:34461597759(synchronize)和 34461599480(labeled);并发组在任何基准开始前取消了 synchronize 运行;34461599480 为最终扫描运行
  • 范围:变更日志条目对应的完整配置矩阵:5 个基准点(并发 1/2/4/8/16,TP16/DCP16,每点 4 节点)以及每点默认的 kimi_tool_call_schema 评测。无裁剪或评测选择修饰。

结果

  • 10:20 UTC:基准点 c1/c2/c4/c8 自 09:40 UTC 起并行运行;c16 与五个仅评测任务排队等待。目前无失败。
  • 11:00 UTC:c1/c2/c4/c8 仍在运行(已 80 分钟,处于模型加载后的 60 分钟采样窗口内);c16 与评测仍在排队。
  • 11:35 UTC:基准点 c1/c2/c4/c8 通过;c16 以及 c2/c4/c16 的评测正在运行,c1/c8 的评测排队中。与 2026-08-21 已发布基线的逐点中期对比(配方相同;最终值以入库为准):
并发 每 GPU 总吞吐(tok/s) 输出 tok/s TTFT 中位数(s) ITL 中位数(s) 交互性中位数(tok/s) 采样请求数
1 917.3 → 971.0(+5.8%) 112.3 → 117.4(+4.6%) 2.158 → 1.589(-26.4%) 0.00448 → 0.00428(-4.5%) 223.2 → 233.6(+4.7%) 232 → 238
2 974.0 → 1036.2(+6.4%) 131.4 → 142.9(+8.8%) 2.078 → 1.345(-35.3%) 0.00456 → 0.00434(-4.8%) 219.3 → 230.5(+5.1%) 372 → 390
4 1489.0 → 1553.7(+4.3%) 164.5 → 170.4(+3.6%) 2.070 → 1.408(-32.0%) 0.00479 → 0.00444(-7.3%) 208.8 → 225.3(+7.9%) 507 → 517
8 2161.8 → 2320.8(+7.4%) 220.2 → 244.7(+11.1%) 15.369 → 7.730(-49.7%) 0.00509 → 0.00478(-6.1%) 196.5 → 209.1(+6.4%) 840 → 900

这四个点的无效结果统计:c1 11/252(全部为预热)、c2 24/419(预热 22)、c4 44/568(全部为预热)、c8 68/1000(预热 90)——冒烟中出现的采样阶段“无内容”响应除 c2 的两个外未再出现。

  • 11:55 UTC:最终扫描在评测门槛处失败。 该配置在完整扫描中的默认评测为每点 lm-eval gsm8k(而非裁剪定向运行所用的 kimi_tool_call_schema 冒烟)。已完成的三个评测均未通过 Verify eval scores(strict-match 阈值 0.90):
并发 gsm8k strict(已发布基线 c16 为 0.9727) gsm8k flexible(基线 0.9719) 缺少 #### N 答案行的响应
2 0.8165 0.9090 N/A(未检查样本)
4 0.7991 0.9030 N/A(未检查样本)
16 0.8120 0.9067 219 / 1319(基线运行:0 / 1319)
  • 诊断(首个服务端原因):评测服务健康且 1319 个请求全部完成,但 c16 有 219 个响应在答案行之前于句中截断(响应长度中位数 258 字符,基线样本为 293;无空响应;真正答错 29 个,基线为 36 个)。原因是修复 1/5 采用的 Dynamo 固定版本 ba83080 原生聊天处理器截断响应尾部:其 kimi_k3 XTML 工具解析器会缓冲尾部文本,并在流结束时将其作为“不完整调用”丢弃(dynamo_parsers::tool_calling::xtml::kimi_k3_parser … kimi_k3_incomplete_call,基准前端日志中亦可见),且在 ba83080 上即使请求未声明工具该解析器也会运行。基线使用的 vLLM 聊天处理器路径无法在此镜像上运行(vllm.entrypoints.openai.cli_args 于 2026-08-25 被上游移除,见初始尝试),而 Dynamo main 仍导入同一路径。
  • 基准方面:c1/c2/c4/c8 通过,逐点差值见上表;停止时 c16 基准与 c1/c8 评测仍在运行,由于该 head 上的扫描已无法通过验证,将由 finish 取消。
  • 停止决定:不允许任何运行时补丁,因此剩余的范围内选项是将配方的 Dynamo 固定版本改为仅对声明了工具的请求进行工具解析的提交(fix(frontend): scope tool parsing to request semantics ai-dynamo/dynamo#13359,2026-08-20 合并,或更新版本),并重新验证 worker 与 vLLM 3696c77 的兼容性。这需要新的定向冒烟(约 1.7 小时)加完整扫描(约 3 小时),超出本会话剩余预算,因此该候选在使用 1/5 次修复后以失败结束,分支保留供人工审阅。建议维护者后续:在 agg-gb200-dcp16-dspark4-maxseq2-mooncake-agentic.yaml 中随本镜像一并提升 dynamo.hash / identity.frameworks.dynamo,或等待上游 Dynamo 前端适配新的 vLLM cli_args 位置以恢复 dyn-chat-processor: vllm

@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Stop: final sweep on head 256689714 failed its gsm8k eval gate (strict 0.80–0.82 vs threshold 0.90; baseline 0.9727) because the pinned Dynamo ba83080 native chat processor truncates response tails; the original vLLM chat-processor path cannot run on any vLLM image after 2026-08-25. Details and the recommended Dynamo-pin follow-up are in the Final full sweep comment. Repairs used: 1/5. The remaining sweep jobs are cancelled, the PR returns to draft without sweep labels and is closed; the branch is retained for manual review.


停止: head 256689714 上的最终扫描未通过 gsm8k 评测门槛(strict 0.80–0.82,阈值 0.90;基线 0.9727),原因是固定版本 Dynamo ba83080 的原生聊天处理器会截断响应尾部;原有的 vLLM 聊天处理器路径无法在 2026-08-25 之后的任何 vLLM 镜像上运行。详情及建议的 Dynamo 版本提升后续步骤见最终完整扫描评论。已用修复:1/5。剩余扫描任务将被取消,PR 回到草稿状态并移除扫描标签后关闭;分支保留供人工审阅。

@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Klaud Cold: failed. Finishing cleanup; owned child runs will be stopped and checked before closure.


Klaud Cold:failed。正在完成清理;将先停止并确认自有子运行的状态,再关闭 PR。

@github-actions

Copy link
Copy Markdown
Contributor

@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Klaud Cold: failed. All owned runs are terminal. Repairs: 1. Runs: 34446587513, 34449828534, 34461597759, 34461599480, 34473834826.

PR closed; the exact-candidate branch is retained for manual review. The interruption does not prove image incompatibility.


Klaud Cold:failed。所有自有运行均已结束。修复次数:1。运行:34446587513, 34449828534, 34461597759, 34461599480, 34473834826

PR 已关闭;保留该候选的分支,等待人工审查。运行中断不能证明镜像不兼容。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Development

Successfully merging this pull request may close these issues.

1 participant